Papers with task formulation
TaxFree: a Visualization Tool for Candidate-free Taxonomy Enrichment (2022.aacl-demo)
Copied to clipboard
| Challenge: | In this paper, we present an open source system for taxonomy visualisation and automatic taxonomies enrichment without pre-defined candidates. |
| Approach: | They propose an open source system for taxonomy visualisation and automatic taxonomie enrichment without pre-defined candidates on the example of WordNet-3.0. |
| Outcome: | The proposed system can be used for visualisation and inspection of taxonomies without pre-defined candidates on WordNet-3.0. |
DATE: Detecting Anomalies in Text via Self-Supervision of Transformers (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent deep learning methods for anomalies in images learn better features of normality in an end-to-end self-supervised setting. |
| Approach: | They propose to use a novel pretext task to learn a deep learning model for Anomaly Detection in text to train a model to discriminate between different transformations applied to visual data. |
| Outcome: | The proposed method outperforms state-of-the-art methods on 20Newsgroups and AG News datasets in the semi-supervised setting and in the unsupervised setting. |
Preventing Critical Scoring Errors in Short Answer Scoring with Confidence Estimation (2020.acl-srw)
Copied to clipboard
Hiroaki Funayama, Shota Sasaki, Yuichiroh Matsubayashi, Tomoya Mizumoto, Jun Suzuki, Masato Mita, Kentaro Inui
| Challenge: | Recent Short Answer Scoring systems use Quadratic Weighted Kappa (QWK) but it is unsatisfactory when measuring their effectiveness in actual usage. |
| Approach: | They propose a task formulation of Short Answer Scoring (SAS) that matches actual usage and extracts as many scoring predictions that are not critical scoring errors (CSEs). |
| Outcome: | The proposed system predicts scores with zero critical scoring errors (CSEs) for 50% of test data at maximum by filtering out low-reliability predictions on the basis of a certain confidence estimation. |
Measuring the Effect of Influential Messages on Varying Personas (2023.acl-short)
Copied to clipboard
| Challenge: | a new task estimates the response a persona might have upon seeing a news message . a first benchmark dataset is used to evaluate the performance of the proposed task . |
| Approach: | They propose a task to estimate the response a persona might have upon seeing a news message. |
| Outcome: | The proposed task estimates the response a persona might have upon seeing a news message. |
Narrative Question Answering with Cutting-Edge Open-Domain QA Techniques: A Comprehensive Study (2021.tacl-1)
Copied to clipboard
| Challenge: | Recent advances in open-domain question answering (ODQA) have led to human-level performance on many datasets. |
| Approach: | They provide a comprehensive and quantitative analysis about the difficulty of book QA . they compare the results of their research with extensive ODQA experiments . |
| Outcome: | The proposed model outperforms existing models on event-oriented questions on the NarrativeQA dataset. |
S2SPMN: A Simple and Effective Framework for Response Generation with Relevant Information (D18-1)
Copied to clipboard
| Challenge: | Existing work on how to generate relevant and informative responses is focusing on how dialogue systems generate information from large dialogue corpus. |
| Approach: | They propose to use dialogue corpus to generate relevant responses by using prototypes to extract semantic information from PMN. |
| Outcome: | The proposed model outperforms classical and strong baseline models in generating relevant and informative responses. |
Challenges in Information-Seeking QA: Unanswerable Questions and Paragraph Retrieval (2021.acl-long)
Copied to clipboard
| Challenge: | Existing pretrained language models have solved reading comprehension benchmarks, but datasets with information-seeking queries remain challenging. |
| Approach: | They analyze why answering information-seeking queries is more challenging . they manually annotate 800 unanswerable examples across six languages . |
| Outcome: | The proposed model outperforms human annotators on 800 unanswerable examples across six languages. |
FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework (2022.acl-long)
Copied to clipboard
Santiago Castro, Ruoyao Wang, Pingxuan Huang, Ian Stewart, Oana Ignat, Nan Liu, Jonathan Stroud, Rada Mihalcea
| Challenge: | Existing video understanding evaluation frameworks that use fill-in-the-blanks do not reflect real-world tasks. |
| Approach: | They propose to use fill-in-the-blanks as a video understanding evaluation framework and introduce a novel dataset that collects multiple perspectives on the same video. |
| Outcome: | The proposed framework does not share the weaknesses of the current state-of-the-art language-informed video understanding tasks, namely: (1) video question answering using multiple-choice questions, where models perform relatively well because they exploit linguistic biases in the task formulation; (2) video captioning, which relies on an open-ended evaluation framework that is often inaccurate because system answers may be perceived as incorrect if they differ in form from the ground truth. |
OLMES: A Standard for Language Model Evaluations (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing models claim to perform better on tasks measuring model capabilities, but there is no standard setup for reproducible evaluations. |
| Approach: | They propose a document that is documented and practical for reproducible LLM evaluations and includes recommendations from existing literature and new experiments. |
| Outcome: | The proposed standard identifies and reviews the varying factors in evaluation practices adopted by the community, such as prompt formatting, choice of in-context examples, probability normalizations, and task formulation. |
COLA: Contextualized Commonsense Causal Reasoning from the Causal Inference Perspective (2023.acl-long)
Copied to clipboard
Zhaowei Wang, Quyet V. Do, Hongming Zhang, Jiayao Zhang, Weiqi Wang, Tianqing Fang, Yangqiu Song, Ginny Wong, Simon See
| Challenge: | Existing efforts to detect commonsense causation from the causal inference perspective are inadequate to seize commonsensical causations. |
| Approach: | They propose a task to detect commonsense causation between two events in context . they propose 'contextualized commons sense causal reasoning' framework that uses covariates to remove confounding effects . |
| Outcome: | The proposed framework can detect commonsense causality more accurately than baselines. |
SOCCER: An Information-Sparse Discourse State Tracking Collection in the Sports Commentary Domain (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods for state tracking are limited and state changes are less densely distributed over utterances. |
| Approach: | They propose to turn to simplified, fully observable systems that show some of these properties. |
| Outcome: | The proposed system shows that state changes occur infrequently while messages are "chatter" it allows for rich descriptions of state while avoiding the complexities of other settings. |
ForecastQA: A Question Answering Challenge for Event Forecasting with Temporal Text Data (2021.acl-long)
Copied to clipboard
| Challenge: | Existing automated forecasting studies rely on structured data to predict future events. |
| Approach: | They propose a question-answering task that limits access to unstructured text data . they use a crowdsourced dataset to form a restricted-domain, multiple-choice, question-announcement task . |
| Outcome: | The proposed model achieves 61.0% accuracy on the dataset, which still lags behind human performance by about 19%. |
Cross-lingual Contextualized Phrase Retrieval (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Phrase-level dense retrieval has shown many appealing characteristics in downstream NLP tasks. |
| Approach: | They propose a task formulation of dense retrieval, cross-lingual contextualized phrase retrieval . they extract pairs of cross-linguistic phrases using word alignment information . |
| Outcome: | The proposed task formulation surpasses baselines on the phrase retrieval task and a downstream task, i.e., machine translation, and achieves top-1 accuracy 13 points higher. |
Adaptive Text Anonymization: Learning Privacy-Utility Trade-offs via Prompt Optimization (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for anonymizing textual documents lack flexibility to adapt to diverse requirements. |
| Approach: | They propose a task formulation in which anonymization strategies are automatically adapted to specific privacy–utility requirements. |
| Outcome: | The proposed framework achieves better privacy–utility trade-off than existing baselines on open-source language models while remaining computationally efficient and effective on larger closed-source models. |
Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing tools for examining and fixing missing captions are lacking in mobile UIs. |
| Approach: | They propose a task for automatically generating language descriptions for UI elements from multimodal input including both the image and structural representations of user interfaces. |
| Outcome: | The proposed task can generate captions from image and structural representations of UI elements. |
A Dataset for Tracking Entities in Open Domain Procedural Text (2020.emnlp-main)
Copied to clipboard
Niket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson, Eduard Hovy
| Challenge: | Existing tasks require only a small set of attributes to track state changes in procedural text. |
| Approach: | They propose a task where given a procedural text as input, the task is to generate a set of state change tuples for each step. |
| Outcome: | The proposed task generates state change tuples from a set of pre-defined attributes for each step and predicts them from an open vocabulary. |
Come hither or go away? Recognising pre-electoral coalition signals in the news (2021.emnlp-main)
Copied to clipboard
Ines Rehbein, Simone Paolo Ponzetto, Anna Adendorf, Oke Bahnsen, Lukas Stoetzer, Heiner Stuckenschmidt
| Challenge: | In this paper, we decompose the task of recognizing from the news coverage leading up to an election the (un)willingness of political parties to form a coalition into two related, but distinct tasks. |
| Approach: | They propose a task of recognizing from news coverage the (un)willingness of political parties to form a coalition from text and a sub-task of predicting the polarity of the signal. |
| Outcome: | The proposed approach improves over a strong monolingual transfer learning baseline. |
Goal-Driven Explainable Clustering via Language Descriptions (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing formulations neither consider the users’ goals nor explain clusters’ meanings. |
| Approach: | They propose a task formulation that represents both the goal and the explanations as free-form language descriptions. |
| Outcome: | The proposed method produces more accurate and goal-related explanations than previous methods. |
Don’t Just Say “I don’t know”! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies investigate ways to refuse to answer unknown questions . Large Language Models (LLMs) display a significant level of overconfidence when answering questions that they are aware of. |
| Approach: | They propose a self-alignment method to utilize Large Language Models to enhance its response-ability to unknown questions. |
| Outcome: | The proposed method is superior to baseline methods on four types of unknown questions. |
Dual-Stage Multi-Task Syntax-Oriented Pre-Training for Syntactically Controlled Paraphrase Generation (2024.findings-acl)
Copied to clipboard
| Challenge: | Syntactically controlled paraphrase generation (SCPG) aims to generate sentences with syntactic structures resembling given exemplars. |
| Approach: | They propose a dual-stage multi-task pre-training scheme that uses a series of structure-oriented and syntax-oriented tasks to generate sentences with syntactic structures resembling given exemplars. |
| Outcome: | The proposed method outperforms existing methods on all possible variants of SCPG tasks and significantly outperformed the popular T5 model. |
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Empathy plays a pivotal role in fostering prosocial behavior, often triggered by the sharing of personal experiences through narratives. |
| Approach: | They propose to use contrastive learning with masked LMs and supervised fine-tuning with large language models to improve empathy understanding in NLP models. |
| Outcome: | The proposed methods show that there is low agreement among annotators and that cultural differences are a factor in their interpretation of empathy. |